Skip to main content

Clinical AI

Clinical AI is machine learning applied where the output influences care for an individual patient. That is a different activity from analytics: it carries regulatory obligations, safety requirements and accountability that operational dashboards do not.


Application areas​

  • Triage and risk stratification — early warning scores, deterioration prediction, prioritising who is seen first
  • Diagnostic support — suggesting or ranking differential diagnoses
  • Imaging — detection and measurement; see imaging AI
  • Documentation — ambient capture and summarisation of clinical encounters
  • Coding support — proposing ICD and SNOMED CT codes from notes
  • Medication safety — interaction and dosing checks
  • Screening — expanding reach where specialist capacity is scarce

Regulation​

Software whose purpose is diagnosis, prevention, monitoring, prediction, prognosis or treatment is generally a medical device — Software as a Medical Device (SaMD) — regardless of whether it is embedded in hardware.

That means, depending on jurisdiction:

  • Risk classification proportional to potential harm
  • A quality management system (typically ISO 13485)
  • Clinical evaluation evidence
  • Post-market surveillance and incident reporting
  • Change control: retraining a model is a change to a regulated product

Jurisdictions differ (FDA, EU MDR/AI Act, and national regulators), and a tool that is unregulated in one setting may be a device in another. Establish the regulatory position before building, not after piloting.


Safety requirements​

Clinician in the loop, meaningfully. Automation bias is real: people accept confident-looking suggestions. Presenting a prediction as advice does not transfer accountability if the interface makes disagreement difficult.

Explainability proportional to risk. Clinicians need enough of the basis to judge whether the output makes sense for this patient.

Calibrated uncertainty. Confident wrong answers are the dangerous failure mode. Abstention — "not enough information" — is a legitimate and often underused output.

Fail-safe behaviour. What happens when the model is unavailable, or receives inputs outside its training distribution?

Monitoring in production. Performance decays as practice, populations and upstream data change. Silent degradation is the norm without instrumentation.

Documented scope. Which population, which setting, which inputs — and, explicitly, where it must not be used.


Honest questions before deployment​

  1. What decision changes because of this output? If none, do not deploy it.
  2. What is the current standard of care, and does the model beat it?
  3. Who is accountable when it is wrong?
  4. Was it validated on a population like the one it will serve?
  5. How will a clinician override it, and is that path as easy as accepting?
  6. What is the plan for retraining, revalidation and withdrawal?